Papers with low-resource language scenarios
A Semantic Uncertainty Sampling Strategy for Back-Translation in Low-Resources Neural Machine Translation (2025.acl-srw)
Copied to clipboard
Yepai Jia, Yatu Ji, Xiang Xue, Shilei@imufe.edu.cn Shilei@imufe.edu.cn, Qing-Dao-Er-Ji Ren, Nier Wu, Na Liu, Chen Zhao, Fu Liu
| Challenge: | Back-translation methods rely on large-scale parallel corpora to enhance performance, but ignore the semantic quality of monolingual data. |
| Approach: | They propose a method which prioritizes sentences with higher semantic uncertainty as training samples by computationally evaluating the complexity of unannotated monolingual data. |
| Outcome: | The proposed method improves translation accuracy and fluency by +1.7 on all three translation tasks. |
A Multilingual Topic Model for Learning Weighted Topic Links Across Corpora with Low Comparability (D19-1)
Copied to clipboard
| Challenge: | Existing models implicitly assume that documents in different languages are highly comparable, a false assumption. |
| Approach: | They propose a multilingual topic model that learns weighted topic links and connects cross-lingual topics only when the dominant words defining them are similar. |
| Outcome: | The proposed model outperforms existing models in low-resource language tasks and outperformed LDA and previous models in classification tasks using documents’ topic posteriors as features. |